AI Learning Series · Part 23

Agent Harness Architectures

The model is one call; the harness is everything around it that assembles context and executes the loop. This is the architecture survey: how production agents are actually structured, and why every design choice traces back to the token-economics and memory-bandwidth walls of docs 06–09.

Software Context Solutions (09)
→
Agents End-to-End (12)
→
KV Cache Types (22)
→
Harness Architectures
→
Prefill & Decoding (24)

01 The Big Picture

An LLM API call is stateless, single-shot, and blind: it sees exactly the tokens you send and does exactly one pass. Everything that makes an agent — persistence, tools, planning, memory, permissions — is software wrapped around that call. That software is the harness.

Doc 09 catalogued the parts: skills, rules, workflows, memory banks, subagents. Doc 12 animated the tool-use loop. This doc zooms out to the architecture level and asks the engineering question: given the hardware constraints, why are harnesses shaped the way they are? The answer keeps coming back to three walls established in docs 06–09:

Tokens are currency

Every token in the prompt is re-sent and re-priced every turn (doc 07). A harness that lets context grow linearly with the task pays a compounding bill.

Cache economics rule structure

Cached prefix ≈ 0.1× cost. Harnesses are literally shaped so the expensive parts stay prefix-stable (doc 07) — that's why skills load on demand instead of sitting in the system prompt.

Long context is slow and dilute

Every decode step reads the whole KV cache (docs 06, 08, 22), and attention quality degrades on long transcripts (doc 19, lost-in-the-middle). Small per-call context isn't just cheaper — it's smarter.

Keep the mapping discipline from doc 09: for every mechanism below, we name the bottleneck it relieves. A harness feature that relieves no bottleneck is decoration.

02 What a Harness Is — Precisely

A harness is everything around the model that (a) assembles the context sent to it and (b) executes the loop that turns its outputs into actions. The model proposes; the harness disposes — it runs tools, spawns workers, writes memory, enforces permissions, and decides what the model sees next turn. Production harnesses (Claude Code is the canonical example) decompose into six layers. Read top to bottom as "what the model sees" → "what runs around it":

layer 1 · conversationTranscript in the context window — the only state the model natively sees layer 2 · memoryFiles + database outside the window: memory bank, decisions, playbooks layer 3 · tools & subagentsExecutors the model can invoke — including fresh-context workers layer 4 · skills & rulesProgressive-disclosure knowledge + behavioral constraints layer 5 · context engineeringThe compiler that assembles each turn's prompt from the layers above layer 6 · control loopPermissions, worker management, continuation, budgets, termination

Layer 1 is inside the model; layers 2–6 are the harness proper. The single-agent context loop — system prompt + tool schemas + memory files + transcript, re-sent every turn — works precisely because of doc 07's prefix caching: the stable opening (system prompt, tools, skill index) is prefilled once at the cached rate, and each turn only prefills the delta. Compartmentalization keeps that prefix stable; stability keeps it cheap.

03 Why Harnesses Look Like This

The context window is the scarce resource — not disk, not network, not even model quality. Three pressures dictate the architecture:

Transcript growth vs. quality. Raw transcripts grow O(turns × tool-output size). Past roughly the mid-window, attention dilution sets in — the model increasingly misses information buried in the middle (doc 19). A 200K window is not a 200K scratchpad; effective capacity is far smaller. The harness must therefore treat the window as a working set, not a database.
Caching rules ordering. Anything that changes every turn must sit after the cache-stable prefix, or the whole suffix re-prefills at 10× the price (doc 07). This is why harnesses inject timestamps, search results, and user messages at the end, and keep tool schemas and skill descriptions in fixed order at the front.
Output tokens are 3–5× input. Verbose tool outputs pasted into the transcript are the worst of both worlds: expensive fresh input once, then cached-but-still-attention-consuming forever after. Compressing at the source — "return a summary, not the file" — beats compressing after the fact.
🧭
The design principle in one line: the orchestrator's context should scale with task complexity, not with total work done. Everything else — subagents, memory banks, skills, summarization — is a mechanism for enforcing that separation.

04 How It Runs — One Harness Turn, Seven Steps

Step through a single orchestrator turn of a Claude Code–style harness and watch where tokens go — and where they don't.

ORCHESTRATOR — one persistent context system prompt + tools + skill index — CACHED PREFIX (0.1×) session state + loaded skill bodies + memory excerpts transcript — append-only: user msg → turns → summaries ↑ assembled by the context compiler each turn (cache-stable first, volatile last) ⚡ one model call over this context CONTROL LOOP — permission gate · worker mgmt · continuation tool call: read_file(src/auth.py) → allowed? → execute → result routed WORKER SUBAGENT FRESH context per task task brief ≤ tapered budget burns its own tokens, dies after MEMORY BANK (disk) decisions · lessons · playbooks — persists beyond the window brief summary only (~0.5K tok) write loop: continue / finish

The key accounting trick is step 6: the worker may have burned 30K tokens exploring a codebase, but the orchestrator's transcript only grows by the ~0.5K summary. That is doc 09's "compress via tool" pattern, realized as a process boundary: the orchestrator pays for complexity, workers pay for volume.

05 Context Assembly Math

Context assembly is a compiler: it takes source (memory bank, session state, skill files, transcript) and emits one token stream per request, optimizing for the cache. The emitted prompt has three regions:

// the per-turn compilation, in emission order (doc 07 rules) static prefix = system prompt + tool schemas + skill index // cache-stable across the session session state = memory excerpts + loaded skill bodies + plan // stable within a task turn-local = latest results + summaries + user msg // fresh every request

Each region bills at a different rate (doc 07's table), so the cost of a turn is:

cost/request ≈ cached_prefix × 0.1 + fresh_input × 1.0 + output × 3

Worked example. Orchestrator: 12K stable prefix, 2K fresh per turn, 0.5K output per turn:

with cache: 12,000×0.1 + 2,000×1.0 + 500×3 = 1,200 + 2,000 + 1,500 = 4,700 units/turn without cache: 14,000×1.0 + 500×3 = 15,500 units/turn // ~3.3× worse, before TTFT pain

Growth rates. A naive single agent that pastes every tool output into one transcript accumulates T turns of unbounded context: per-turn cost grows linearly and total session cost grows O(T²) even with caching (the transcript itself is the growing prefix). Worse, attention quality decays as the transcript passes the effective window (doc 19). The subagent-constrained architecture bounds per-call context: orchestrator stays O(task complexity), each worker starts fresh with a tapered budget (≤ some threshold, e.g. 30–50K tokens), and peak context per call stays flat no matter how long the task runs.

⚖️
The trade, made explicit: subagents add coordination overhead — N workers means N full prompt re-reads of the stable prefix. You win when exploration volume is large relative to decision density: reading 30K tokens to produce a 0.5K summary is a 60:1 compression; running that exploration in the orchestrator's transcript would tax every subsequent turn forever.

06 Architecture Patterns Compared

Beyond the orchestrator–worker core, harnesses differ in who holds context and who decides. Anthropic's agent-building guidance names three archetypes — orchestrator–workers, planner–executor, and evaluator–optimizer — plus the decentralized extreme:

PatternContext costCoordination overheadWhen to use
Single agent + toolsOne transcript; grows with total work — O(T²) session costZero — but all exploration tax lands in one windowShort, focused tasks where tool outputs are small and summarizable
Orchestrator–worker (Claude Code–style)Orchestrator O(task complexity); workers pay their own volumeN subagent calls; briefs must be written carefullyParallelizable, read-heavy work: research, multi-file changes, fan-out search
Planner–executorPlanner context stays small (plan, not transcripts); executor per-stepPlan revision loop between the two rolesLong-horizon tasks where sequencing mistakes are expensive
Evaluator–optimizerTwo contexts: generator + critic; critic re-reads output each roundR rounds × re-evaluation costTasks with a checkable quality bar: tests, specs, rubrics
Decentralized swarmMany small contexts — minimal per-agent bloatHighest: agents communicate peer-to-peer; no global viewEmerging/experimental; robustness over efficiency
BlackboardShared structured store; each specialist reads only its sliceRead/write policy design on the blackboardMulti-expert problems with no fixed pipeline order (classic AI: HEARSAY-II)

Planner–executor in one sentence

The planner never touches raw tool output; it produces and revises a plan artifact. Executors turn plan steps into actions. The plan is the interface — a natural compression point, since a plan is O(steps), not O(observations).

Blackboard in harness terms

A shared file or DB the orchestrator and workers both read/write — decisions, findings, open questions. It substitutes for shared context: agents coordinate through state on disk, which is cheap to store and selectively read, instead of tokens in a window.

07 The Bottleneck Mapping Table

The mapping discipline from doc 09, extended to the full harness. Every mechanism earns its complexity by relieving a specific bottleneck from docs 06–09:

MechanismBottleneck it relievesHardware/cost link
Cache-stable static prefixRepeated prefill compute of identical opening tokens0.1× cached rate (doc 07); TTFT collapse
Subagent fresh contextsTranscript growth taxing every later decodeKV-cache read per decode step (docs 06, 22)
Summaries as return valuesRaw tool output entering the persistent transcriptFresh-input + permanent attention tax
Tapered worker budgetsUnbounded single-context growth → dilutionLost-in-the-middle (doc 19)
Memory bank (files/DB)Window as working set, not databaseDisk is ~10⁵ cheaper than context tokens
Skills: name+desc in contextStatic knowledge inflating the cached prefixPays 0.1× on the index, 1.0× only on demand
RAG / hybrid retrievalRelevance filtering before token spendEmbedding lookup ≪ prefill of wrong docs (doc 11)
Tool permission gateBlast radius of injected instructionsDefense-in-depth (doc 16)
Write policies on memoryWrite cost vs. read value asymmetryStore decisions, not transcripts

08 Memory, Skills, Retrieval — The Persistence Layer

Memory bank & procedural memory

Playbooks ("how we deploy here"), decisions, and lessons-learned live in files the harness reads into context when relevant. This is procedural memory: not what happened, but how to act. It survives session death because it lives on disk, priced in bytes, not tokens.

Episodic & comparative memory

What got persisted this session, and why. The write policy is an economics problem: raw transcripts have huge write cost and near-zero read value; decisions and lessons have tiny write cost and high read value. Persist the delta, not the log.

RAG inside the harness

Hybrid retrieval (embeddings + keyword, doc 11, doc 19) decides which memory enters context. Retrieval is a token-spending filter: pay embeddings to avoid prefilling the wrong 10K tokens.

Skills = progressive disclosure, priced

A skill ships a name + one-line description in the cached prefix (a few dozen tokens each) and a detailed body on disk. The model sees the menu; the harness loads the dish only when the task matches. The cost math: 200 skills × ~30 tokens of description ≈ 6K tokens of menu at the 0.1× cached rate — versus megabytes of documentation that would otherwise be pasted, prefilled, and attention-taxed every turn. If a skill triggers, its body (k tokens) is paid at full rate once, on the turns that need it.

09 Failure Modes

Context rot & attention dilution

Long transcripts don't just cost more — they degrade. Mid-context information is systematically under-attended (doc 19). Symptom: the agent "forgets" an instruction stated 80K tokens ago even though it's still in the window. The fix is architectural: shorter effective contexts, summaries, re-injection of key constraints near the turn-local region.

Tool-call storms

A loop of failed searches or retry storms burns output tokens at 3–5× while adding junk to the transcript. Harnesses need budgets: max calls per task, loop detection, and hard termination — the control-loop layer's real job.

Prompt injection from tool output

Every tool result is untrusted input inside the context. A malicious web page read by a worker can command the orchestrator. Defenses are layered — permission gates, output quarantining, least-privilege tools (doc 16) — because context assembly cannot distinguish "text" from "instructions."

Coordination overhead blowup

Spawning N workers for a task one agent could do re-reads the stable prefix N times and serializes on brief-writing. Subagents are a compression bet; if the subtask's exploration volume is small, the bet loses. Measure: worker tokens ÷ summary tokens should be ≫ 1.

10 Mental Models

Harness = OS kernel

Context window = RAM (scarce, fast, priced per byte); memory bank = disk (cheap, persistent, needs explicit loads); subagents = processes (own address space, isolated, return exit codes); skills = dynamically linked libraries (mapped on demand); the control loop = the scheduler with a permission ring. Lets you reason about: why "the model" is the CPU and the harness is everything that makes a computer out of it.

An OS preempts processes; subagents can't be interrupted mid-call — the harness only chooses when to spawn and what to accept back.
Context assembly = a compiler

Source files (memory, skills, plan) + symbol table (session state) → one emitted object file (the request). The cache is the compiler's incremental-build cache: touch nothing in the prefix and the build is nearly free. Lets you reason about: prompt ordering as a compilation constraint, not a style choice (doc 07).

Compilers optimize for correctness; the context compiler optimizes for cache-prefix stability, which occasionally conflicts with what's "cleanest" to include.
Manager who never reads the files

A good orchestrator delegates: it writes a precise brief, gets a summary back, and makes the decision. It never reads the 30-page report — not because it can't, but because reading it would clog its one working memory forever. Lets you reason about: O(task complexity) vs. O(total work) contexts.

Managers also need taste for the details occasionally — some tasks warrant the orchestrator reading raw output, at a deliberate context cost.

11 Anti-patterns

✓ Do

Bound the orchestrator's transcript with summaries and worker isolation; write decisions & lessons to the memory bank; keep the static prefix genuinely static; give workers tapered budgets and demand compressed returns.

✗ Don't

Run the god-agent: one transcript, everything pasted in, "the window is 200K so who cares." That design pays O(T²) session cost, hits attention dilution, and rots — while still billing you at full price every turn.

✓ Do

Treat memory writes as investments with a read-value test: "will a future turn plausibly load this?" Persist the decision and the reason, not the transcript.

✗ Don't

Dynamically inject volatile data into the cached prefix — timestamps, reordered tools, refreshed memory files mid-session. One changed token invalidates the entire suffix at 10× price (doc 07's fragility rule).

12 Misconceptions & Closing Insights

"More agents = better agent." Agents don't add intelligence; they add context isolation and parallelism — both of which cost coordination. A swarm that shares nothing reproduces work; an orchestrator with lazy briefs gets garbage summaries. Pattern choice follows the workload's volume-to-decision ratio.

"Subagents save tokens." They usually spend more total tokens (N fresh contexts, N prefix re-reads). What they save is the orchestrator's peak context — which is what actually compounds across turns and degrades quality. Cost-per-call down, sometimes total cost up; know which you're optimizing.

"Memory = a bigger context window." Long-context models raise the ceiling but not the economics: decode still reads the whole KV cache (docs 06, 22) and attention still dilutes (doc 19). The harness's memory hierarchy exists precisely because brute-force window growth is the wrong shape of solution.

"The harness is boilerplate around the model." Invert it: the model call is one opcode; the harness is the machine. Permissions, context assembly, memory policy, and worker scheduling determine cost, safety, and capability far more than the raw model choice at fixed quality tiers.

🧠
Closing insight — the paradigm shift, named: static code shipped instructions to a deterministic machine; a harness ships context to a probabilistic one. The six layers of section 02 are the new "architecture diagram" of software: conversation, memory, tools, skills, the context compiler, and the control loop. Engineering this stack — not prompt-crafting alone — is what doc 03 called context engineering, now seen whole.
🗺️
Next: the harness controls what goes into the model; doc 14 covered what comes out token-by-token, doc 24 covers how providers schedule the very compute your harness is renting, doc 26 asks what happens when adversarial content rides in through the tools and memory this architecture connects, and Perception · AI as an Observability Stack covers how to instrument the harness in production — the links below point into the rest of the series.